Papers with synthesis quality
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding (2025.acl-demo)
Copied to clipboard
| Challenge: | Experimental evaluations show RT-VC delivers a 13.3% reduction in latency . voice conversion modifies speech to match the timbre of a target speaker while preserving content information. |
| Approach: | They propose a zero-shot real-time voice conversion system that leverages an articulatory feature space to naturally disentangle content and speaker characteristics. |
| Outcome: | The proposed system achieves a CPU latency of 61.4 ms, representing a 13.3% reduction in latency. |
TCSinger: Zero-Shot Singing Voice Synthesis with Style Transfer and Multi-Level Style Control (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing models fail to generate singing voices rich in stylistic nuances for unseen singers due to multifaceted nature of singing styles. |
| Approach: | They propose a zero-shot SVS model for style transfer across cross-lingual speech and singing styles and multi-level style control. |
| Outcome: | Experimental results show that TCSinger outperforms baseline models in synthesis quality, singer similarity, and style controllability. |
Empowering Diffusion Models on the Embedding Space for Text Generation (2024.naacl-long)
Copied to clipboard
| Challenge: | Recent work adapts diffusion models to textual data by diffusing on the embedding space. |
| Approach: | They propose an embedding diffusion model based on Transformer to solve the problem of embeddable space and denoising model. |
| Outcome: | The proposed model is more efficient than previous methods on seminal text generation tasks and is superior to existing models. |
SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing retrieval-augmented generation methods overlook citation graph structure, adapt poorly to complex queries, and yield fragmented, hard-to-verify syntheses. |
| Approach: | They propose a retrieval-augmented generation framework that addresses these gaps by combining adaptive retrieval and symbolic reasoning. |
| Outcome: | Extensive experiments show that SciRAG outperforms prior systems in factual accuracy and synthesis quality. |
ProsodyFlow: High-fidelity Text-to-Speech through Conditional Flow Matching and Prosody Modeling with Large Speech Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Text-to-speech (TTS) models have been developed to generate high-quality speech. |
| Approach: | They propose an end-to-end TTS model that integrates large self-supervised speech models and conditional flow matching to model prosodic features effectively. |
| Outcome: | The proposed model improves synthesis quality and efficiency compared to existing models, showing that it generates more prosodic and expressive speech synthesizing. |
MeanAudio: Fast and Faithful Text-to-Audio Generation with Mean Flows (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in Text-to-Audio Generation (TTA) systems suffer from slow inference speed, authors report . authors demonstrate that MeanAudia achieves state-of-the-art performance in single-step audio generation . |
| Approach: | They propose a text-to-audio generator capable of rendering realistic sound with only one function evaluation. |
| Outcome: | The proposed system achieves state-of-the-art performance in single-step audio generation. |
TCSinger 2: Customizable Multilingual Zero-shot Singing Voice Synthesis (2025.findings-acl)
Copied to clipboard
Yu Zhang, Wenxiang Guo, Changhao Pan, Dongyu Yao, Zhiyuan Zhu, Ziyue Jiang, Yuhan Wang, Tao Jin, Zhou Zhao
| Challenge: | Existing zero-shot singing voice synthesis models depend on phoneme and note boundary annotations, limiting their robustness and producing poor transitions between phonemes and notes. |
| Approach: | They propose a multi-task multilingual zero-shot SVS model with style transfer and style control based on various prompts. |
| Outcome: | Experimental results show that TCSinger 2 outperforms baseline models in subjective and objective metrics across multiple related tasks. |